Papers with human ratings

2 papers
RankME: Reliable Human Ratings for Natural Language Generation (N18-2)

Copied to clipboard

Challenge: Existing studies have shown that human evaluation for natural language generation often suffers from inconsistent user ratings.
Approach: They propose a rank-based magnitude estimation method which combines continuous scales and relative assessments to improve the reliability of human ratings.
Outcome: The proposed method significantly improves the reliability and consistency of human ratings compared to traditional evaluation methods.
A Dual-Perspective NLG Meta-Evaluation Framework with Automatic Benchmark and Better Interpretability (2025.acl-long)

Copied to clipboard

Challenge: Existing evaluation metrics are insufficient to meet requirements for natural language generation.
Approach: They propose a dual-perspective NLG meta-evaluation framework that focuses on different evaluation capabilities and a method of automatically constructing benchmarks without requiring new human annotations.
Outcome: The proposed framework improves interpretability and provides better performance for 16 representative LLMs.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations